Papers with learnt word embeddings
Time-Aware Word Embeddings for Three Lebanese News Archives (2020.lrec-1)
Copied to clipboard
| Challenge: | a large corpus of newspaper archives has been generated, but historians have struggled to analyze it manually for decades. |
| Approach: | They propose to train word embeddings from three large Lebanese news archives, which collectively consist of 609,386 scanned newspaper images and span 151 years. |
| Outcome: | The embeddings are trained using a Google Tesseract 4.0 OCR engine and a benchmark of analogy tasks to evaluate their accuracy. |